Repository navigation
feat(recipe): add Qwen3.5 text MXFP8 long-context SFT - #6300
Merged
cuichenx merged 5 commits intoOct 7, 2026
Merged
Conversation
Signed-off-by: Chen Cui <chcui@nvidia.com>
Contributor
|
Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing: Add |
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
kamran-nvidia
approved these changes
Oct 7, 2026
ilml
added a commit
to ilml/Megatron-Bridge
that referenced
this pull request
Oct 8, 2026
Resolve conflicts with the Qwen3.5 text MXFP8 long-context SFT recipe from main (NVIDIA-NeMo#6300). Both sides registered a new recipe in recipes/qwen/__init__.py and recipes/qwen/gb200/__init__.py, so both exports are kept. The qwen35.py import block takes main's imports, which include everything this branch needs. Signed-off-by: Tom Long <tolong@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a Qwen3.5-35B-A3B text-only 128K MXFP8 SFT recipe for 16 GB200 GPUs, with recipe exports and focused tests. The recipe uses TP1/CP8/EP16, one MTP layer, packed CoderForge data, MXFP8 parameter gather and gradient-buffer reuse, CuTeDSL grouped MLP, single grouped weights, and HybridEP. Model and dataset revisions are pinned.
The recipe starts directly from
_sft_common()and retains workload-specific settings. Redundant assignments now inherit the common builder, model provider, dataset, and precision defaults.Validation: 105 focused text/Qwen recipe tests and all-file pre-commit pass on this refactor. Before/after normalized configurations match with offline HF metadata and the real model provider. Independent subagent review found no blocking issues. This PR contains only the new recipe, exports, and focused tests; no verification card or documentation changes.
Historical launcher-versus-Bridge comparisons completed 100 updates with matching losses, but used an uncommitted recipe precursor and Megatron-LM PR #7611. They are not clean-checkout verification of this PR. Masked MTP correctness requires NVIDIA/Megatron-LM#7611 or an equivalent fix in the Bridge MCore pin; this PR does not update dependencies.